Papers with crowdsourcing platform
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)
Copied to clipboard
| Challenge: | Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues. |
| Approach: | They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms. |
| Outcome: | The proposed corpus includes parallel text and speech data of 21 Japanese dialects. |
Crowdsourcing-based Annotation of the Accounting Registers of the Italian Comedy (L18-1)
Copied to clipboard
Adeline Granet, Benjamin Hervy, Geoffrey Roman-Jimenez, Marouane Hachicha, Emmanuel Morin, Harold Mouchère, Solen Quiniou, Guillaume Raschia, Françoise Rubellin, Christian Viard-Gaudin
| Challenge: | CIRESFI project aims to reassess a theatrical heritage that has often been considered inferior to that of the two major, royally-privileged theaters. |
| Approach: | They propose a double annotation system for new handwritten historical documents . crowdsourcing platform is set up to perform labeling and transcription of the documents based on budget data . |
| Outcome: | The proposed system is based on a database of 25,250 pages of registers of the Italian Comedy of the 18th century. |
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)
Copied to clipboard
Vilelmini Sosoni, Katia Lida Kermanidis, Maria Stasimioti, Thanasis Naskos, Eirini Takoulidou, Menno van Zaanen, Sheila Castilho, Panayota Georgakopoulou, Valia Kordoni, Markus Egg
| Challenge: | a large corpus of online content has been developed via large-scale crowdsourcing. |
| Approach: | They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform. |
| Outcome: | The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines. |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |
Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French (L18-1)
Copied to clipboard
| Challenge: | a crowdsourcing project in the field of Swiss German dialects and Swiss French accents collects linguistic data. |
| Approach: | a gamified crowdsourcing platform was set up to collect linguistic data on Swiss German and Swiss French accents. |
| Outcome: | a gamified crowdsourcing platform collects linguistic data on Swiss German and Swiss French accents . the platform has provided 470,000 localizations, with 7,500 registered users and 30,000 anonymous visitors . |
Nonsense!: Quality Control via Two-Step Reason Selection for Annotating Local Acceptability and Related Attributes in News Editorials (D19-1)
Copied to clipboard
| Challenge: | Annotation quality control is critical for building reliable corpora through linguistic annotation. |
| Approach: | They propose a method to control annotation quality using two-step reason selection using a crowdsourcing platform. |
| Outcome: | The proposed method retains the annotations with satisfactory quality out of the entire annotations mixed with those of low quality. |
Rigor Mortis: Annotating MWEs with a Gamified Platform (2020.lrec-1)
Copied to clipboard
| Challenge: | gamification of the platform should be improved, in order to attract and retain more players. |
| Approach: | They propose to use a gamified crowdsourcing platform to evaluate the intuition of speakers and then train them to annotate multi-word expressions in French corpora. |
| Outcome: | The proposed platform evaluates the speakers' intuition and trains them to annotate multi-word expressions in French corpora. |
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)
Copied to clipboard
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, Joni Salminen
| Challenge: | Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals. |
| Approach: | They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments . |
| Outcome: | The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube. |